Papers with process reward model

7 papers
Progressive Multimodal Reasoning via Active Retrieval (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to improve multimodal large language models' reasoning performance are limited.
Approach: They propose a framework to progressively improve multimodal reasoning capabilities . they propose active retrieval and Monte Carlo tree search to improve MLLMs' reasoning .
Outcome: The proposed framework improves multimodal reasoning capabilities in multimodal large language models.
Boosting Policy and Process Reward Models with Monte Carlo Tree Search in Open-Domain QA (2025.findings-acl)

Copied to clipboard

Challenge: Experimental results show that our approach can effectively improve the performance of both the policy model and the reward model.
Approach: They propose to use Monte Carlo Tree Search for both policy model improvement and reward model improvement to bridge it to more subtle open-domain question answering.
Outcome: The proposed approach surpasses existing methods for annotation and training data with fewer data points and achieves better performance in test-time scaling strategies.
AgentPro: Enhancing LLM Agents with Automated Process Supervision (2025.emnlp-main)

Copied to clipboard

Challenge: Existing frameworks lack explicit supervision during the reasoning process, which may lead to error propagation across reasoning chains.
Approach: They propose a framework which automates process supervision for large language model agents by automatically generating step-level annotations and developing a process reward model based on these annotations.
Outcome: The proposed framework outperforms existing agent-based methods on four datasets and achieves a 6.32% increase in accuracy.
Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple Actions (2026.acl-long)

Copied to clipboard

Challenge: Existing reinforcement learning methods are expensive due to high latency and sample inefficiency . Currently, RL is limited to one-to-one state-action pairs .
Approach: They propose a framework that shifts the training paradigm to Single State Multiple Actions and introduce a group-wise advantage estimator based on the averaged critic outputs.
Outcome: The proposed framework achieves 7.5% and 8.3% success rate improvements on AndroidLab and AndroidWorld over UI-TARS-1.5-7B and attains 1.4x higher training efficiency than existing methods.
How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark (2025.emnlp-main)

Copied to clipboard

Challenge: Prior work has not explored the mechanisms underlying this sensitivity.
Approach: They propose a synthetic benchmark to evaluate Large Language Models’ reasoning robustness against systematically controlled irrelevant context (IC).
Outcome: The proposed model improves in-distribution and out-of-disttribution scenarios while training with strong distractors.
ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling (2026.acl-long)

Copied to clipboard

Challenge: Existing research on inference scaling focuses on unstructured output generation tasks, such as mathematical problems.
Approach: They propose an inference-scaling framework that combines fine-grained beam search with ToolPRM, a process reward model scoring each intra-call decision.
Outcome: The proposed framework outperforms outcome and coarse-grained reward models in predictive accuracy and yields consistent test-time gains on multiple function-calling benchmarks.
Verified Critical Step Optimization for LLM Agents (2026.findings-acl)

Copied to clipboard

Challenge: Critical Step Optimization (CSO) focuses preference learning on verified critical steps where alternative actions demonstrably flip task outcomes from failure to success.
Approach: They propose a method which focuses preference learning on verified critical steps where alternative actions demonstrably flip task outcomes from failure to success.
Outcome: The proposed method outperforms the existing methods on GAIA-Text-103 and XBench-DeepSearch while requiring supervision at only 16% of trajectory steps.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations